Papers with cross-lingual evaluation

6 papers
NollySenti: Leveraging Transfer Learning and Machine Translation for Nigerian Movie Sentiment Classification (2023.acl-short)

Copied to clipboard

Challenge: Africa has over 2000 indigenous languages but they are under-represented in NLP research due to lack of datasets.
Approach: They propose to use a dataset to classify sentiments for cross-domain adaptation for Nigerian and other African languages.
Outcome: The proposed dataset compares the performance of cross-domain adaptation from Twitter domain and cross-lingual adaptation from English domain.
Lower Perplexity is Not Always Human-Like (2021.acl-long)

Copied to clipboard

Challenge: Existing efforts to build human-like computational models have focused on English . a cross-lingual evaluation is needed to build such models, but current research has focused on Japanese .
Approach: They re-examine an established generalization that lower perplexity is not always human-like in Japanese . they propose a cross-lingual evaluation to build human-type computational models .
Outcome: The proposed model lacks universality and lower perplexity is not always human-like . the results suggest a cross-lingual evaluation will be necessary to build human-type models .
LAReQA: Language-Agnostic Answer Retrieval from a Multilingual Pool (2020.emnlp-main)

Copied to clipboard

Challenge: LAReQA tests for “strong” cross-lingual alignment, requiring semantically related cross-language pairs to be closer in representation space than unrelated same-language pair.
Approach: They propose a new benchmark for language-agnostic answer retrieval from a multilingual candidate pool that tests for "strong" cross-lingual alignment . they augment training data via machine translation and find that model performance is improved by augmenting training data through machine translation .
Outcome: The proposed task is based on multilingual BERT (mBERT) and XLM-R.
AM2iCo: Evaluating Word Meaning in Context across Low-Resource Languages with Adversarial Examples (2021.emnlp-main)

Copied to clipboard

Challenge: Existing multilingual evaluation datasets that evaluate lexical semantics "in-context" have various limitations, including limited coverage of high-resource languages and superficial cues.
Approach: They propose to use a set of pretrained language models to evaluate lexical semantics in context.
Outcome: The proposed set shows that current models lag behind human performance in interpreting word meaning in cross-lingual contexts.
Cross-Lingual Auto Evaluation for Assessing Multilingual LLMs (2025.acl-long)

Copied to clipboard

Challenge: Evaluating machine-generated text remains a challenge in NLP for non-English languages . current evaluation frameworks focus on English, revealing a gap in multilingual evaluations .
Approach: They propose a cross-lingual auto evaluation framework that includes evaluator LLMs and a test set specifically designed for multilingual evaluation.
Outcome: The proposed model aligns more closely with human judgments than proprietary models on non-English language evaluations.
Read the Room, Read the Image: Understanding Indirect Speech Acts in Multimodal Visual Contexts (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks focus on explicit context, but do not address context-dependent pragmatic understanding.
Approach: They propose a benchmark for evaluating ISA understanding through integrated reasoning over visual context and dialogue.
Outcome: Experiments show that state-of-the-art models struggle with visually grounded indirect speech acts . linguistic meaning emerges through the relationship between an utterance and situational context .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations